Papers by Zhen Hao Wong
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration (2025.acl-long)
Copied to clipboard
Tianyi Bai, Ling Yang, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Chi Zhang, Lijun Wu, Qiu Jiantao, Wentao Zhang, Binhang Yuan, Conghui He
| Challenge: | Efficient data selection is crucial to accelerate the pretraining of language models . limited research has addressed the inherent conflicts between data selection methods . |
| Approach: | They propose a multi-actor collaborative data selection mechanism that prioritizes data based on its specific criterion and updates prioritization rules using the current state of the model. |
| Outcome: | The proposed model accelerates convergence in LM pretraining and achieves an average relative performance gain of 10.5% across multiple language model benchmarks. |
Data-Centric Perspectives on Agentic Retrieval-Augmented Generation: A Survey (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel at natural language understanding and generation, yet rely on static pre-training data. |
| Approach: | They propose to augment Large Language Models with external retrieval to ground model outputs . traditional RAG is constrained by a fixed retrieve-then-generate routine . authors aim to guide creation of high-quality datasets for next generation of adaptive LLM agents . |
| Outcome: | The proposed model can decompose tasks, issue exploratory queries, and refine evidence through iterative retrieval. |
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks provide only final questions and answers, while lacking intermediate hop-level questions that gradually connect atomic questions to the final multi-hop query. |
| Approach: | They propose to build a multi-hop reasoning model that is primarily constructed automatically by large language models and designed to support step-by-step validation. |
| Outcome: | The proposed benchmark spans multiple domains, contains 1,305 data points, and has no overlap with existing mainstream benchmarks. |